Back

Applications in Plant Sciences

Wiley

Preprints posted in the last 30 days, ranked by how well they match Applications in Plant Sciences's content profile, based on 23 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit.

1
Rclade: automated taxonomic collapsing and geological-timescale annotation of time-calibrated phylogenetic trees in R

Zeng, Z.; Wang, Y.

2026-09-01 bioinformatics 10.64898/2026.08.27.747462 medRxiv
Top 0.1%
12.2%
Show abstract

Background: Reproducible taxonomic collapsing and geological-timescale annotation of time-calibrated phylogenetic trees in R often require coordination among several packages and repeated code for label parsing, clade validation, plotting, and export. Workflow-managed analyses additionally benefit from non-interactive configuration, predictable diagnostics, and machine-readable exit status. Results: We present Rclade, an R package that consolidates the multi-package coordination required for taxonomic collapsing into a streamlined, single-function interface. Rclade provides (1) custom ggproto objects (GeomPolygonStraight/GeomSegmentStraight) that bypass coord_munch() interpolation to achieve straight-edge rendering of collapsed triangles in circular layouts; (2) automatic detection and parsing of four taxonomic-label formats (GTDB, Silva, NCBI, embedded) plus user-supplied custom regex, with explicit input-validation contracts and parsing-accuracy evaluation on real and derived test sets; and (3) workflow embeddability through YAML configuration, library-mode APIs, and standard Unix exit codes. Benchmarks on synthetic and real datasets (200-10,000 synthetic tips and real reference trees up to 10,122 tips; 5 replicates at every scale under a unified fully rendered measurement protocol) show that the full-pipeline overhead is modest for interactive use (median {approx}0.87 s in-session rendering and {approx}8.4 s process-level wall-clock at 10,000 tips). Conclusions: Rclade is a convenience layer over the ggtree/deeptime ecosystem that reduces boilerplate while adding targeted technical improvements for circular-layout rendering and format heterogeneity management.

2
Utilising nuclear encoded plastid DNA to identify donors of grass-to-grass lateral gene transfer

Bourne, N. G.; Payne, L.; Manzi, S.; Besnard, G.; Vorontsova, M. S.; Jobson, R. W.; Chomicki, G. S.; Dunning, L. T.

2026-08-29 evolutionary biology 10.64898/2026.08.26.747220 medRxiv
Top 0.1%
11.3%
Show abstract

Determining the correct donor species/lineages of grass-to-grass lateral gene transfer (LGT) is vital for deducing specific donor features that could help inform the mechanism of transfer. This requires a dataset spanning a broad range of species to achieve the phylogenetic resolution necessary for precise donor inference. As grass-to-grass LGT often involves the transfer of multi-gene DNA fragments, they can contain additional sequences that allow for accurate orthologous comparisons, such as nuclear DNA of plastid origin (NUPTs). Here we systematically scan for NUPTs in the genomes of four Alloteropsis semialata accessions, whose LGTs have previously been characterised. Using the abundant Panicoideae chloroplast sequences, we reconstruct NUPT phylogenies and infer two lateral acquisitions: one from Paniceae/Digitaria and another from Andropogoneae/Eremochloa adjacent to a previously identified LGT. We then assembled and included an additional 12 Eremochloa chloroplast genomes in the analysis and showed the likely donor was Eremochloa attenuata. Subsequent short-read mapping from E. attenuata to the nuclear region flanking this NUPT showed consistent coverage across the region, including the previously identified LGT, supporting co-transfer. Overall this study highlights the potential for NUPTs to better identify the donors of grass-to-grass LGT.

3
PlantOmicsGWAS: An end-to-end, reproducible framework for plant genome-wide association and genomic prediction using linear and pan-genome references

Khan, F. S.; Yassin, A.; Rehman, S. u.; Sun, T.; Wang, X.; Sun, H.; Abe-Kanoh, N.; Su, Y. H.; Guo, L.; Ye, W.

2026-08-20 bioinformatics 10.64898/2026.08.16.745120 medRxiv
Top 0.1%
10.6%
Show abstract

Genome-wide association studies (GWAS) play a crucial role in unraveling the genetic foundations of complex traits in plants but are also hampered by the application of heterogeneous tools, incompatible file formats and disparate computational environments. Existing GWAS frameworks are often restricted to a single linear reference genome, limiting the capacity for the analysis of structural variations and presence/absence variations (PAV) within plant populations. These issues pose obstacles to reproducibility, scalability, and comprehensive investigations. Here, we present PlantOmicsGWAS, an open-source Python framework for reproducible plant genome-wide association analysis and genomic prediction. It integrates reference indexing, FASTQ quality control, alignment, variant calling, VCF normalization, PLINK conversion, linkage disequilibrium analysis, population-structure estimation, association testing, marker scoring, genomic prediction, and visualization within a unified Linux and HPC workflow. The framework supports conventional linear-reference analyses and includes an optional pangenome-oriented module for working with multiple assemblies and graph-derived variation. Using a Vitis benchmark dataset containing 120 accessions and 118,247 graph-derived variants, PlantOmicsGWAS reduced manual workflow fragmentation and generated standardized association outputs. This tool provides a modular and extensible platform for plant GWAS and pan-GWAS workflows while retaining compatibility with established command-line tools and common genotype formats. The GWAS workflow described herein is adaptable to a range of sequencing methods and plant genomes, bridging research on crop related issues across various biological levels, from the individual organism to entire populations. PlantOmicsGWAS implements Bayesian sparse linear mixed modeling (BSLMM) through GEMMA for multi-trait association discovery, while also supporting FaST-LMM, regression-based approaches, and machine-learning algorithms (Random Forest, XGBoost) as benchmarking alternatives. The PlantOmicsGWAS, a versatile toolkit is available at GitHub https://github.com/plantomicsgwas1-boop/PlantOmicsGwas_V1 and on Linux and HPC platform (https://pypi.org/project/PlantOmicsGwas/1.0.2/).

4
Introducing entropy-based metrics for quantifying edge- and macro-shape complexity in leaves and beyond

Trauden, T.; Rakotomalala, A. A. N. A.; Junker, R. R.; Sauressig, L.; Trauden, K.; Munoz Andres, M.; Dannoritzer, R.; Farwig, N.; Pinkert, S.

2026-08-27 ecology 10.64898/2026.08.26.747315 medRxiv
Top 0.1%
6.6%
Show abstract

Leaf shape is a fundamental trait of plant ecological strategies, influencing biotic interactions and ecosystem functioning. However, established quantitative metrics fail to capture subtle variations and irregularities, require user-based reference points or are challenging to compare among taxa with broadly different leaf shapes. In addition, established metrics typically conflate (aggregate) leaf edge complexity and macro-shape complexity, despite their independent functional significance and genetic foundations. Here, we introduce an entropy-based framework to quantify two new complexity metrics: edge complexity and macro-shape complexity. Based on three case studies, we show that these metrics outperform aggregate metrics in predicting Quercus robur chemical traits, provide more intuitive interspecific classifications, and strongly align with human perception. In addition, edge and macro-shape complexity show high complementarity, while aggregate metrics are highly redundant and typically strongly related to leaf area. Emerging as the strongest predictor of leaf chemistry and key visual cue for complexity as perceived by humans, the effects of edge complexity highlight the under-appreciated functional significance of leaf margins. Our framework and the proposed entropy-based complexity metrics thus promise to help unlock the potential of growing digital image archives of leaves, including images from herbaria and fossils, and are technically readily applicable to shapes of algae, bacteria, pollen, and beyond. The accompanying package ShapeComplexity enables the broad application of entropy-based metrics, providing a powerful tool to explore how the shape of organisms and biological structures influences ecological strategies, biotic interactions, and ecosystem functioning while tracking spatial and temporal variation.

5
Global vascular plants reveal persistent gaps across taxa and ecoregions

Maciel, E. A.

2026-08-28 evolutionary biology 10.64898/2026.08.24.746674 medRxiv
Top 0.1%
6.3%
Show abstract

Biodiversity aggregators such as GBIF provide unprecedented access to global biodiversity data, yet their representativeness remains uneven across space and taxa. This study examined the spatial and taxonomic structure of global vascular plant data available on GBIF. Six filters were applied to the GBIF vascular plant dataset, resulting in the removal of 54% of all records. Together, the filters explained more than 90% of the identified spatial issues, with duplicate and missing coordinates accounting for most of the variation. A higher number of occurrence records was associated with a greater number of spatial issues. Record distributions became progressively more even at finer taxonomic levels, from orders to species. The time series of occurrences for species, genera, and families increased sharply after 1800 and continued to rise, with no apparent stabilisation. Of the 824 ecoregions covered, 73 accounted for 72% of all occurrence records. These ecoregions spanned all continents but were strongly concentrated in Europe, followed by North America and Oceania. The analyses reveal four key patterns: (1) data volume is positively associated with spatial issues; (2) a small number of taxa account for a large proportion of records, whereas many are represented by relatively few; (3) occurrence data aggregated by GBIF have increased continuously since 1800; and (4) record coverage remains highly uneven across the world's ecoregions. These results highlight the substantial contribution of biodiversity data aggregators to expanding access to biological information while demonstrating the persistent spatial and taxonomic biases that shape their contents. Such biases should be explicitly considered when assessing data completeness and quality and when using aggregated occurrence records to infer global biodiversity patterns.

6
PlantAI: A Multi-Agent System for Plant Functional Genomics Analysis and Biological Knowledge Interpretation

Wu, T.; Yang, Z.; Shi, J.; Zou, M.; Wu, Y.; Jiang, S.; Xia, C.; Kong, L.; Yang, L.; Xia, Z.

2026-08-18 bioinformatics 10.64898/2026.08.14.744760 medRxiv
Top 0.1%
6.0%
Show abstract

Plant functional genomics requires the integration of sequence, expression, evolutionary, regulatory and literature evidence. However, the corresponding analyses are often distributed across disparate programs, scripts and databases, creating substantial barriers to task organization and result interpretation. Here, we present PlantAI, a multi-agent system that integrates bioinformatics analysis, project-level process tracking and knowledge-assisted interpretation. A Main Agent coordinates two complementary routes: an analysis route that invokes bioinformatics tools for RNA-seq and gene-family analyses, and a knowledge route that uses PlantAI-RAG for knowledge retrieval and evidence synthesis. PlantAI-RAG currently contains 31,207 plant-science literature records, comprising approximately 3.82 million normalized entities and 8.25 million literature-supported relation assertions. In an evaluation using plant-science questions, it achieved a Gold evidence-assertion recall of 86.7%, while strict accuracy ranged from 77% to 82% across three independent evaluator models. We further demonstrate an end-to-end task using 24 rice RNA-seq libraries collected under salt stress, spanning transcriptome analysis, candidate-family screening, HXK/HKL family analysis and knowledge-assisted interpretation, and prioritize OsHXK8 for experimental validation. By preserving analysis artifacts, run manifests, logs and environment records, PlantAI supports result verification and repeat execution while linking project-derived results to traceable literature evidence. Together, these capabilities provide an integrated and auditable framework to support plant functional genomics research.

7
Chloroplast Genome Evolution, Heteroplasmy, and Inverted Repeat Dynamics in the Elymus Complex (Triticeae, Poaceae): Insights from Single-Molecule Sequencing of Elymus ciliaris and Comparative Analysis of St-Genome Lineages

Karimi, N.; Zhang, Y.; Saeidi, H.; Schwarzacher, T.; Liu, Q.; Heslop-Harrison, J. S.

2026-08-07 genomics 10.64898/2026.08.03.742459 medRxiv
Top 0.1%
4.3%
Show abstract

Background/ObjectivesElymus sensu lato (Poaceae) is arguably the largest and most complex genus in the tribe Triticeae. It includes hybrids and polyploids based on x=7 chromosomes, all including the St genome, forming a valuable genepool for forage grass and cereal breeding. Analysis of chloroplast genome diversity and structural dynamics is critical for resolving maternal lineages, reticulate evolution and biodiversity across this agronomically important complex, refining their taxonomy, conservation and exploitation. MethodsWe sequenced the complete chloroplast genome (plastome) of Elymus ciliaris (4x=2n=28; StStYY genome composition) using ultra-long Oxford Nanopore single-molecule reads and compared it to 76 additional chloroplast genomes representing major St-genome lineages in Elymus s.l. (Pseudoroegneria St; Elymus s.s. StH, StY; Thinopyrum StJ/E; Campeiostachys StYH; Kengyilia StYP). We analyzed structure, nucleotide diversity, inverted repeat (IR) dynamics, and phylogenetic signal. ResultsThe E. ciliaris chloroplast genome was 135,004 bp long (38.3% GC) with a canonical quadripartite structure. Single-molecule reads (n=74) revealed heteroplasmy: two Small-Single-Copy (SSC) orientations at 30%:70% frequency, indicating an inversion polymorphism. Across the Elymus group, comparative analysis of chloroplast assemblies showed high structural conservation but lineage-specific IR-boundary shifts. Kengyilia exhibits exceptional IR expansion. Nucleotide diversity hotspots localize to the large single-copy region, especially in StY lineages. Phylogenies recover a monophyletic St-containing clade but do not delineate genera, reflecting reticulate evolution, with North American/Southeast Asian and Eurasian geographic sub-clades. ConclusionsSingle-molecule sequencing uncovered heteroplasmy with an inversion polymorphism in a single plant of Elymus ciliaris, hidden in short read assemblies. There were no other polymorphisms, as expected for chloroplast sequences (except for technical homopolymer variation). Our analyses showed that a Pseudoroegneria-like St chloroplast genome predominates as the maternal donor across Elymus polyploids. Variable regions and IR dynamics offer strong models for chloroplast genome evolution in reticulate lineages and suggest exploiting plastome variation to complement nuclear biodiversity studies.

8
High-Molecular-Weight Genomic DNA Extraction from Recalcitrant Australian Plants: An Optimised CTAB Protocol for Anigozanthos

Rajput, R.; Saha, L.; Ahmed, Z.; Naiker, P.; Do, L.; Bisset, A.; Hooper, C.

2026-08-31 plant biology 10.64898/2026.08.29.741951 medRxiv
Top 0.1%
3.3%
Show abstract

High-phenolic plant genera present a major technical limitation in genomic research. Standard extraction approaches that perform reliably across diverse flora often perform poorly when applied to recalcitrant taxa, producing low DNA yield and integrity incompatible with sequencing requirements. The genus Anigozanthos (Kangaroo paws) from the family Haemodoraceae exemplifies this problem. We identified key physicochemical factors governing extraction failure in this genus and resolved them through targeted modifications to lysis chemistry and contaminant management. The resulting protocol achieved a near threefold improvement in DNA purity, substantially reducing contaminant carry over and consistently yielded high-integrity, long DNA fragments (DIN > 7) across a diverse sample set spanning cultivated and wild material across four diverse genera of Haemodoraceae. We also tested a straightforward purity assessment framework that can be implemented in any standard molecular laboratory, enabling rapid pre-submission quality assessment without the need for specialised equipment. Together these advances open a practical path to genomic characterisation of Anigozanthos that establishes a transferable model for genomic research across Australia ' s chemically complex native flora.

9
Chromosome-level genome assemblies and annotations of Amaranthus spinosus, Amaranthus acanthochiton, Amaranthus arenicola, and Amaranthus floridanus

Raiyemo, D. A.; Werle Noe, I.; Kaur, R.; Whitt, L.; Carey, S. B.; Hale, H.; Lewis, K. J.; Womack, L.; Harkess, A.; Llaca, V.; Fengler, K.; Patterson, E. L.; Gaines, T. A.; Tranel, P. J.

2026-08-25 genomics 10.64898/2026.08.21.746229 medRxiv
Top 0.1%
3.1%
Show abstract

Amaranthus L. spans aggressive agricultural weeds, ornamentals, and ancient pseudocereals. Species within the genus vary in morphology, environmental tolerance, and sexual systems, making them well-suited for studying reproductive evolution and plant adaptation. To investigate sex chromosome architecture within the genus, we generated chromosome-level assemblies of a monoecious amaranth (Amaranthus spinosus) and three dioecious species (A. acanthochiton, A. arenicola, and A. floridanus) using PacBio high-fidelity (HiFi) long reads. We paired these data with Dovetail Genomics Omni-C sequencing to achieve haplotype resolution for A. spinosus and A. acanthochiton, and we used reference-guided scaffolding for the remaining two species. The assemblies are highly contiguous, with sizes ranging from 394.24 to 607.10 Mbp, contig N50 from 0.63 to 8.76 Mbp, and scaffold N50 from 22.44 to 37.97 Mbp. Evaluation of the assemblies and annotations revealed 96.3 to 97.6%, and 97.6 to 98.3% BUSCO completeness, respectively. Comparative genomic analysis revealed that the Chromosome 1 inversions and Robertsonian fusion previously reported in A. tuberculatus are conserved in A. acanthochiton and consistent with the architecture of A. arenicola and A. floridanus, suggesting that the evolution of dioecy in this clade predates subsequent speciation. In parallel, multiple homologs of Rf1 on Chromosome 3 of A. spinosus, a monoecious species that exhibits spatial separation of male and female flowers and is closely related to the dioecious A. palmeri, were identified. Together, this study provides foundational resources for advancing evolutionary, ecological, and agronomic research across the genus, including herbicide resistance evolution and weediness traits.

10
MIRA: an open source and user-friendly software to automate counting and sizing of fungal spores

Mejias, J.; Adreit, H.; Blanc, A.; Lubin, N.; Jolivet, C.; Guyot, V.; Brayle, O.; Poncelet, N.; Fournier, E.; Wicker, E. P.; Carlier, J.; Tharreau, D.; Ravel, S.

2026-08-07 plant biology 10.64898/2026.08.06.743221 medRxiv
Top 0.2%
2.5%
Show abstract

BackgroundThe quantification of fungal spores constitutes a fundamental metric in phytopathology, serving as the primary variable for inoculum standardization and being used as a proxy for disease severity. Historically, spore quantification has relied on manual hemocytometry, which remains the most precise counting process to date, where chambers such as the Malassez slide are used to count a subsample of the inoculum. However, this method applied manually is highly labor-intensive, time-consuming, and can be prone to operator-dependent variability. To overcome these limitations, we introduce MIRA (Microscopy Image Recognition & Analysis), a novel open-source software integrating You Only Look Once (YOLO) deep learning algorithms. Featuring a user-friendly graphical interface, MIRA is adaptable to multiple camera systems and supports advanced object detection models, including YOLOv11 and YOLOv26. ResultsWe demonstrate that MIRA can be used to accurately detect and count spores from several phytopathogenic fungi, automatically measure spore surface area, and to differentiate spores across different genera. In an exhaustive comparative analysis using Pyricularia oryzae spores as an example, MIRA was benchmarked against manual gold-standard counting slides (Malassez and Kova) and indirect spectrophotometric methods (SPARK). The P. oryzae model loaded via MIRA achieved a strong correlation (R = 0.96) with manual gold standards while reducing processing time by over 90% for high-concentration samples (10 spores/mL). Beyond this benchmark, we also successfully tested specific YOLO models designed to recognize macro- and microconidia of Fusarium oxysporum f. sp. cubense, a model for Pseudocercospora fijiensis, and a single multiclass model capable of identifying six different rice pathogenic fungi. We provide comprehensive tutorials for operating the software and training custom detection models for free using Roboflow and Google Colab. MIRA is available both as open-source Python code and as standalone executables for Windows and Linux. ConclusionsMIRA provides a rapid, accurate, and highly reproducible alternative to manual spore counting, effectively removing a major bottleneck in phytopathology workflows. By combining advanced YOLO-based deep learning with an accessible interface and comprehensive training resources, MIRA makes accessible automated image analysis for researchers without programming expertise. Moreover, MIRA drastically improves the efficiency of high-throughput disease phenotyping and can be adapted for a wide range of microscopic quantification tasks across various biological disciplines.

11
An R-Based Adaptive Quadtree Spatial Tiling Workflow for Boundary-Exact GBIF Species Occurrence Mining within User-Defined KML Polygons

Pradhan, P.

2026-08-20 ecology 10.64898/2026.08.16.745083 medRxiv
Top 0.2%
2.4%
Show abstract

Global Biodiversity Information Facility (GBIF) occurrence retrievals for an irregularly shaped region are limited by the API spatial query capabilities - rectangular envelopes or size/vertex-limited WKT polygons - neither of which conform to protected areas, sacred groves, wetlands, panchayat or municipal boundaries or any other arbitrary KML polygon of interest queried by users. This paper presents and validates an open, self-contained, adaptive spatial-tiling protocol that (i) ingests any KML polygon of any shape, size and location on earth, breaks it into a set of GBIF API-compatible rectangular tiles, (ii) queries, cleans and clips the individual records to the target polygon, and (iii) summarises the inventory with a generic diversity-completeness-rarefaction module, with minimal manual re-parameterisation between sites. The protocol implements an iterative quadtree refinement algorithm that adapts tile number, size and location to the target polygon geometry, is combined with a fault-tolerant pagination/retry query system, a boundary-exact two-step clipping procedure and a Chao1-based completeness assessment to ensure statistical comparability between sites of different spatial extent and sampling intensity. The algorithm is implemented in open R source (sf, terra, rgbif, tidyverse) with the tiling algorithm controlled by the four parameters only (initial cell size, area floor, tile overlap threshold, recursion limit), with default settings on a new site by simply changing the input file path. This paper describes in detail its five main components - (i) polygon input and validation, (ii) quadtree adaptive tiling, (iii) polygon coverage verification, (iv) tile-wise GBIF query with retry/shrink pagination and partial data retention, (v) boundary-exact deduplication, clipping and diversity estimation. A downstream generic module estimates diversity, Chao1 richness/completeness and Hurlbert rarefaction, for each taxonomic rank and generates rank-ordered diversity tables as output. The generalisability of algorithm to multiple sites has been demonstrated with second polygon (Sonamukhi Sal forest dominated stretch, Bankura district, West Bengal; approx. 610 sq km) that differs from the first (Bishnupur Sal forest dominated stretch; 938 sq km) in both size and complexity (10 vs 34 KML vertices) and report the tiling and diversity metrics comparable results across the two polygons. With no parameter changes, the algorithm generated 135 adaptive query tiles for Sal forest dominated stretch adjoining Bishnupur, and 86 tiles for Sal forest dominated stretch Sonamukhi SDFP, covering completely the area of both polygons. The number of tiles per 100 sq km is comparable between the two runs (14.4 vs 14.1 tiles) despite the 35% difference in polygon size and 3.4x vertex count. The tile-wise querying with retry/shrink pagination retrieved 6,169 GBIF records (excluding errors) with boundary-exact clipping across 404 species for Bishnupur and 1,222 GBIF records (excluding errors) across 271 species for Sonamukhi; the generic diversity module processed the records without further parameter changes and generated comparable metrics for each rank at both sites. The protocol addresses a general bioinformatic challenge in polygon-based GBIF queries, is provided as an open, reusable, documented method which has been validated on two sites. Because the protocol has so far been validated on only two polygons that differ markedly in size, shape and observer regime, it may be regarded as an initial cross-site validation rather than a comprehensive benchmark, and recommend testing on a broader, globally distributed set of polygons before the approach is treated as a general-purpose standard.

12
ChlORIS: Chloroplast Orthologs Resource & Identification Suite

Tong, Y.; Rossetto Marcelino, V.; Turnbull, R. B.; Verbruggen, H.

2026-08-11 genomics 10.64898/2026.08.05.743164 medRxiv
Top 0.2%
2.2%
Show abstract

Chloroplast or plastid genomes are essential resources for studying the evolution and diversity of algae and land plants. Although thousands of plastid genomes have been sequenced, their full potential has not been realised; derived resources such as orthogroup databases and reference datasets for metagenomic profiling remain underdeveloped. We present the ChlORIS database to address these problems across all algal phyla. From 2,254 publicly available algal plastid genomes, after dereplication we clustered 2,531 orthogroups from the annotated proteins and selected 496 orthogroups with consistent gene naming, enabling cross-genome comparisons of homologous plastid proteins. We further selected 224 core orthogroups, each containing more than 10 protein sequences, for which we produced score-calibrated hidden Markov models (HMMs), multiple sequence alignments and predicted protein structures. The value of these resources for phylogenomics is demonstrated through a large-scale plastid phylogeny of 859 taxa spanning all major algal lineages. We characterised the protein HMMs by cross-referencing them to Pfam domains and calibrated score cutoffs for reliable detection. The metagenomic database, HMM library, nucleotide and amino acid alignments, predicted structures and protein metadata, cross-linked to UniProt and InterPro (Pfam), are openly available on the ChlORIS website at https://chloris.codeberg.page/.

13
PyiTOL: reproducible Python workflows for iTOL annotation and taxonomic monophyly assessment

Zeng, Z.; Wang, Y.

2026-08-29 bioinformatics 10.64898/2026.08.27.747471 medRxiv
Top 0.2%
2.1%
Show abstract

Motivation: The Interactive Tree of Life (iTOL) is widely used to display and annotate phylogenetic trees, but managing its format-sensitive annotation files impede reproducible high-throughput analyses. Among the maintained Python packages and versions evaluated, none combined template generation, taxonomic monophyly assessment and iTOL batch operations. Results: PyiTOL validates inputs, generates 31 iTOL template schemas (22 accepted by the live batch uploader), performs LCA-based monophyly classification with nested-monophyly detection, sampling-completeness states and polyphyletic subgroup decomposition, plus API upload and session replay. On a topology-constructed benchmark, all calls matched prespecified labels for 4,389 groups; on a 700-genome tree, binary mono/non-mono calls agreed with ETE4 for 409 genera; 17,294 GTDB R232 genera were processed in about 17 s. Availability and Implementation: PyiTOL 1.0.3 (Python [≥]3.10; Linux, macOS and Windows) is MIT-licensed at https://github.com/ZengZichao/PyiTOL and archived with test data at Zenodo (https://doi.org/10.5281/zenodo.22106806).

14
Four numbers, one axis: deep learning models reveal what leaf spectrum constrains about Farquhar-von Caemmerer-Berry photosynthesis

Ray, R.; Maloof, J.; Magney, T.

2026-08-28 plant biology 10.64898/2026.08.27.747677 medRxiv
Top 0.2%
2.0%
Show abstract

Leaf reflectance spectra are emerging as a viable substitute for gas-exchange measurements of photosynthetic capacity, with a community benchmark reporting that a spectrum accurately recovers most Farquhar-von Caemmerer-Berry (FvCB) parameters. This study re-scores the recovery under dataset-blocked, species-blocked, and leave-one-dataset-out designs, measuring the split-half reliability of each curated parameter. We constructed a convolutional encoder that maps a spectrum to the four parameters through a fixed, differentiable FvCB decoder trained on measured assimilation. A conspecific of 97.4% of held-out leaves were present in the training set, and accuracy is lost along the dataset axis but not along the species axis. Under blocked evaluation, a spectrum constrains a single capacity axis. Jmax25 retains only 17% of its recovery when Vcmax25 is held constant, and the Jmax25:Vcmax25 ratio is not predicted above a median null. The curated values of TPU25 are not reproducible, whereas those of Rday25 are well determined, but its recovery fails due to the loss. The published study measures interpolation rather than transfer, and spectra constrain less of the FvCB parameter space than assumed, including the carboxylation to electron transport balance. Routing predictions through explicit biochemistry makes identifiability measurable, although it does not improve prediction accuracy.

15
PhenoStream: A Cyberinfrastructure for Automated and AI-Based Crop Trait Extraction from Aerial Imagery

Varela, S.; Ruhter, J.; Sacks, E.; Zheng, X.; Allen, D.; Hale, A.; Landry, C.; Kuang, X.; Long, B.; Zhu, Y.; Proma, S.; Kaur, S.; Jarquin, D.; Morrison, J.; Leakey, A.

2026-08-30 plant biology 10.64898/2026.08.26.747008 medRxiv
Top 0.2%
1.7%
Show abstract

The integration of digital technologies for high-throughput field phenotyping is critical for accelerating crop improvement in agriculture. However, extracting traits from remote sensing data remains constrained by fragmented workflows, manual intervention, and limited interoperability among existing tools, resulting in delays that hinder timely biological insight and decision-making. To address these challenges, we present PhenoStream (Phenotyping Streaming), a scalable, end-to-end cyberinfrastructure designed to automate the full lifecycle of aerial imagery-based phenotyping, from data acquisition to plot- and genotype-level inference. The framework integrates automated data ingestion from distributed field sites, geospatial processing, and AI-enabled trait extraction within a unified, user-accessible graphical interface. Its modular and extensible architecture supports adaptable trait modeling and seamless integration of new data sources, enabling deployment across diverse crops, environments, and experimental designs. We demonstrate the system across a large multi-location field trial network of bioenergy crops, where it enables high-throughput characterization of spatiotemporal growth dynamics, genotype-by-environment (GxE) interactions, and predictive modeling of key agronomic traits. By significantly reducing processing latency and manual effort, the platform facilitates near-real-time analysis and reproducible workflows. This work establishes a generalizable and scalable pathway for operationalizing very-high-spatial resolution aerial phenotyping in agricultural research. By bridging data acquisition and analytics, the end-to-end cyberinfrastructure provides a foundation for integrating heterogeneous and unstructured data streams--including remote sensing, environmental, and management data--toward data-driven decision making in agriculture.

16
Destructive harvest validation of high-throughput measurements show that water use efficiency is unaffected by moderate drought in tobacco

Stutz, S. S.; Edquilang, R.; Bernacchi, C. J.; Ort, D. R.

2026-08-31 plant biology 10.64898/2026.08.28.747842 medRxiv
Top 0.2%
1.7%
Show abstract

Water-use efficiency (WUE), the ratio of accumulated plant biomass to water lost through transpiration has conventionally been determined using a destructive single-point measurement. Recent advances in high-throughput phenotyping now enable repeated, non-destructive estimation of biomass and WUE. However, these digital measurements must be statistically validated against conventional destructive methods to validate their use as reliable proxies. Therefore, we compared digital biomass determined point clouds produced from multispectral camera scanners with destructive harvests across eight harvests using Samsun tobacco grown under both drought and high-water conditions. WUE efficiency, calculated using the digital biomass estimated from a point cloud and gravimetric water use determinations, were compared to destructive harvest determinations. The coefficient of variation (CV) showed there were no significant differences in digital and destructive measurements for either biomass or WUE. Indicating that digital measurements can be used in place of destructive measurements. Drought plants used significantly less water and were significantly smaller than high-water plants from Harvests 4 through 8. However, there were no significant differences in the ratio of evapotranspiration to leaf area or WUE, indicating that drought plants were simply smaller and used less water than the high-water plants. This work validates that estimating plant biomass from a digital point coupled with continuous gravimetric determination of water use provides a reliable nondestructive measure of WUE in high-throughput measurements across the full plant life cycle.

17
DICAROS: Diffeomorphic Ancestral Shape Reconstruction on Phylogenies

Severinsen, M. L.; Li, J. K.; Lim, W.; Raskin, L. Y.; Yang, G.; Sommer, S.; Hipsley, C. A.; Nielsen, R.

2026-08-22 evolutionary biology 10.64898/2026.08.21.746152 medRxiv
Top 0.2%
1.5%
Show abstract

Reconstructing ancestral morphologies on a phylogenetic tree is a central task in evolutionary morphometrics. Established reconstruction methods, including multivariate Brownian-motion approaches, rely on linear assumptions and do not directly model the correlations between landmarks within a shape, which can oversimplify the reconstructed morphology. The DICAROS method (Diffeomorphic Independent Contrasts for Ancestral Reconstruction of Shapes; Severinsen et al., 2026) instead fuses sibling shapes along branches with large-deformation diffeomorphic (LDDMM) landmark dynamics that model these correlations, so that ancestors remain on the shape manifold. DICAROS was shown to outperform ordinary least-squares, Brownian-motion, and penalized-likelihood reconstruction, particularly on non-symmetric trees. The dicaros package repackages that pipeline as a documented, pip-installable tool that runs on arbitrary landmark datasets from a single command. It handles 2D and 3D landmarks, Newick and NEXUS trees, a choice of Euclidean or Frechet species means, optional anchor-based alignment, and tips backed by a single specimen, and it returns the reconstructed shapes for all nodes together with the tree relabelled at its internal nodes. We demonstrate dicaros on two new datasets: a 2D leaf dataset (217 species) and a 3D guenon skull dataset (22 species).

18
Quantifying Crop Disease Trait Dynamics through Longitudinal Imaging and Temporal Analytics

Ewen, A.; Mendez, R. G.; Al-Shanoon, K.; Omoluabi, D.; Samarasinghe, A.; Oviedo-Ludena, M. A.; Huatatoca, K. C.; Glor, K.; Nabetani, K.; Kutcher, R.; Wang, L.; Stavness, I.; Jin, L.

2026-08-21 bioinformatics 10.64898/2026.08.14.744665 medRxiv
Top 0.3%
1.1%
Show abstract

Reliable and objective phenotyping is essential for plant breeding programs to characterize genetic variation and accelerate crop improvement. Conventional disease assessment relies on expert visual scoring, which is labor-intensive, subjective, and prone to inter- and intra-rater variability. Although image-based phenotyping methods have been proposed, many require manual intervention, specialized imaging setups, or single time-point measurements, limiting their ability to capture disease progression over time. Here, we present a pipeline for longitudinal plant disease phenotyping that quantifies wheat stripe rust and leaf rust progression from time-series images. The pipeline performs semi-automated leaf and automated pustule segmentation from images acquired in situ, enabling objective disease severity estimation with minimal user intervention and without requiring solid backgrounds or manual leaf manipulation or detachment. By extracting temporal traits, including disease severity trajectories and standardized area under the disease progress curve, the method provides a comprehensive characterization of disease development throughout infection. Association between automated and expert assessments was moderate for stripe rust (R2 = 0.58) and strong for leaf rust (R2 = 0.85), while expert inter-rater reliability was moderate for both diseases (ICC = 0.675 and 0.800, respectively). The proposed approach establishes a scalable and reproducible framework for longitudinal disease phenotyping in controlled environments, with broad applications in disease resistance screening and crop breeding.

19
FigTreeKit: A Python toolkit for programmatic FigTree styling, taxonomy-aware clade auditing, and phylogenetic tree rendering

Zeng, Z.; Wang, Y.

2026-08-28 bioinformatics 10.64898/2026.08.27.747475 medRxiv
Top 0.3%
1.1%
Show abstract

FigTree is a long-standing phylogenetic tree viewer, but its GUI-centered workflow does not itself provide a versioned, batch-replayable record of styling operations. We present FigTreeKit, a Python package that serializes a supported subset of FigTree 1.4.4 annotations (!hilight, !color, and !font), audits taxonomy mappings before topology-gated clade collapse, retains selected BEAST-style metadata in the tested fixtures, and invokes a patched FigTree renderer for headless PNG, PDF, and SVG output. Across 60 independently generated balanced trees with 50-10,000 taxa (10 trees per size, each timed 10 times as technical replicates), the tree-level log-log slope of export time was 0.96 (95% confidence interval [CI], 0.91-1.01), which is compatible with approximately linear scaling over the tested range but does not prove it. The 189,801-taxon GTDB R232 bacterial reference tree was parsed and exported as a large-data scalability demonstration. On the 10,122-taxon GTDB R232 archaeal reference tree, the scripted workflow assessed 179 order-level groups; 142 multi-tip groups produced non-trivial collapses, whereas 37 singleton groups did not alter the display. The software is accompanied by 796 passing tests, a golden conformance corpus that includes acceptance tests against the bundled FigTree JAR, deterministic scenario-based topology checks, and an overall statement coverage of 81%, reported as a descriptive engineering metric. FigTreeKit is released under the GPL-2.0-or-later license as the figtreekit package on PyPI, with source code, documentation, and benchmark data archived on Zenodo.

20
From Field Photosynthesis to Genetic Architecture: Insights from the First Dedicated Photosynthesis Hackathon

Matuszynska, A.; Sansa, O.; Adekoya, F. J.; Akinyemi, O. O.; Anokye, E.; Bashir, O. B.; Boyny, Z. Z. F.; Chukwuka, M. K.; Corvest, E.; Dada, A. O.; DellAcqua, M.; Ehemba, G. L.; Finkbeiner, A. J.; Hamabwe, S.; Hodehou, D. A. T.; Kacheyo, O.; Kamfwa, K.; Mhango, K. J.; Abdullahi, W. M.; Munduwe, G.; Ntukidem, S.; Obisesan, O. K.; Odesina, I. S.; Ogechi, N.-U.; Olaoye, O. D.; Olayinka, M. M.; Osei-Bonsu, I.; Rilwan, K. O.; Stival, L.; Tehar, Z.; Tende, R. M.; To, J.; Ugochukwu, U. K.; Unger, A.; van Aalst, M.; Vrbic, D.; Zhang, C.; Theeuwen, T. P. J. M.; Kramer, D. M.; Kromdijk, J.

2026-08-17 plant biology 10.64898/2026.07.24.740625 medRxiv
Top 0.3%
1.1%
Show abstract

Photosynthesis is among the most consequential yet genetically complex traits in crop plants, and translating its natural variation into actionable genomic targets remains a central challenge for breeding climate-resilient varieties. To start addressing this, researchers are generating increasingly large, multi-environment field photosynthesis datasets. Yet, these data have been structurally under-analysed since their inception. Here we report the outcomes of the first dedicated hackathon focused on computational mining of such field data held in Accra, Ghana, in March 2026. Bringing together data scientists, plant physiologists, geneticists, and breeders from Europe and Africa, these interdisciplinary teams used photosynthetic data collected with hand-held fluorometers to genome-wide marker data across four crop species: cowpea (Vigna unguiculata), barley (Hordeum vulgare), common bean (Phaseolus vulgaris), and potato (Solanum tuberosum). Despite using different species and methods, independent teams identified the same three key findings. First, mechanism-informed feature engineering and dynamic modelling recover genetic signals that are not detected or discarded in standard analysis pipelines, resulting in traits with improved heritability and meaningful associations with yield. Secondly, machine learning methods proved effective at uncovering genetic associations, with temporally resolved features substantially outperforming single time-point measurements. Third, raw chlorophyll fluorescence and absorbance traces consistently contained more information and predictive power than the extracted parameters currently used. A defining feature of this event was having experimentalists and data scientists working together, enabling AI approaches to be grounded in domain knowledge and biological mechanisms rather than relying on data alone.